Papers with audio representations
Cross-Lingual Transfer Learning for Speech Translation (2025.naacl-short)
Copied to clipboard
| Challenge: | Increasing interest in building multilingual foundation models for NLP and speech research has led to limited data collection for training ST systems. |
| Approach: | They propose to use Whisper to explore the behavior of multilingual speech foundation models with restricted data. |
| Outcome: | The proposed model can translate to Chinese with a single language, and it can perform transcriptions in other languages. |
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
| Challenge: | We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities. |
| Approach: | They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations. |
| Outcome: | The proposed model outperforms existing models on audio understanding tasks by 1%-84%. |
Towards Cross-Lingual Audio Abuse Detection in Low-Resource Settings with Few-Shot Learning (2025.coling-main)
Copied to clipboard
| Challenge: | Online abusive content detection, particularly in low-resource settings, remains underexplored. |
| Approach: | They propose to use pre-trained audio representations to detect abusive language in Indian languages using Few Shot Learning (FSL) . |
| Outcome: | The proposed model can be used to classify abusive language in 10 languages using the ADIMA dataset with FSL. |
Beyond Transcripts: A Renewed Perspective on Audio Chaptering (2026.acl-long)
Copied to clipboard
| Challenge: | despite its relevance, research on audio chaptering remains limited and predominantly textbased . authors: audio chapterers can't be used linearly because they skim, scrub timelines, jump to relevant moments . acoustic features and learning representations are not used for audio chapterer evaluation . |
| Approach: | They propose to use audio-only architecture to automatically segment audio into coherent sections . they compare audio-based models with acoustic features and a novel audio-oriented architecture . |
| Outcome: | The proposed audio-only architecture outperforms text-based approaches on acoustic features and LLMs. |
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification (2025.naacl-long)
Copied to clipboard
| Challenge: | Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification. |
| Approach: | They propose a training-free method that enhances audio and language representations using mutual feedback. |
| Outcome: | The proposed method outperforms vanilla zero-shot evaluation with significant margins of 0.42%-27.0%. |